Papers by C V Jawahar
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency. |
| Approach: | They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages. |
| Outcome: | The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health). |
IndicSpeech: Text-to-Speech Corpus for Indian Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | India has 22 languages, each of them being spoken by over a million people . the current state of the art text-to-speech systems for Indian languages are lacking in the multimedia domain . |
| Approach: | They propose to train a state-of-the-art TTS system for Hindi, Malayalam and Bengali and publish the results. |
| Outcome: | The proposed system trains neural text-to-speech systems for Hindi, Malayalam and Bengali and makes them publicly available. |
More Parameters? No Thanks! (2021.findings-acl)
Copied to clipboard
| Challenge: | Using network pruning, we find that there are large redundancies in MNMT models. |
| Approach: | They propose a method to prune and retrain redundant parameters of an MNMT model to improve bilingual representations while retaining multilinguality. |
| Outcome: | The proposed method improves bilingual representations while retaining multilinguality. |
CVIT’s submissions to WAT-2019 (D19-52)
Copied to clipboard
| Challenge: | In this paper, we explore multiway-models for Indian languages. |
| Approach: | They propose to use a Transformer architecture to experiment with multilingual models and methods for low-resource languages. |
| Outcome: | The proposed system is feasible in low-resource languages. |